Papers with bilingual model

2 papers
From Pattern to Interpretation. Using Colibri Core to Detect Translation Patterns in the Peshitta. (2022.lrec-1)

Copied to clipboard

Challenge: Using Colibri Core to detect translation patterns in the Hebrew Bible and its Syriac source text is a promising approach, but does not allow the creation of a bilingual model.
Approach: They propose to use Colibri Core to detect n-gram and skipgram patterns in the Hebrew Bible and its Syriac source text.
Outcome: The proposed modeller can detect n-gram and skipgram patterns in both Hebrew and Syriac texts without the need for textual annotations.
Assessing the Role of Data Quality in Training Bilingual Language Models (2025.findings-emnlp)

Copied to clipboard

Challenge: a recent study shows that adding more languages can degrade performance for some languages while improving others.
Approach: They propose a data filtering strategy to select high-quality bilingual training data with only high quality English data.
Outcome: The proposed approach improves bilingual model performance by 2–4% and reduces bilingual models performance gaps to 1%.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations